A Longitudinal Study of Topic Classification on Twitter

نویسندگان

  • Zahra Iman
  • Scott Sanner
  • Mohamed Reda Bouadjenek
  • Lexing Xie
چکیده

Twitter represents a massively distributed information source over a kaleidoscope of topics ranging from social and political events to entertainment and sports news. While recent work has suggested that variations on standard classifiers can be effectively trained as topical filters (Lin, Snow, and Morgan 2011; Yang et al. 2014; Magdy and Elsayed 2014), there remain many open questions about the efficacy of such classification-based filtering approaches. For example, over a year or more after training, how well do such classifiers generalize to future novel topical content, and are such results stable across a range of topics? Furthermore, what features and feature classes are most critical for longterm classifier performance? To answer these questions, we collected a corpus of over 800 million English Tweets via the Twitter streaming API during 2013 and 2014 and learned topic classifiers for 10 diverse themes ranging from social issues to celebrity deaths to the “Iran nuclear deal”. The results of this long-term study of topic classifier performance provide a number of important insights, among them that (1) such classifiers can indeed generalize to novel topical content with high precision over a year or more after training and (2) simple terms and locations are the most informative feature classes (despite training on classes labeled via hashtags). 1 Learning Topical Social Sensors Our objective is to evaluate binary classifiers that can label a previously unseen tweet as topical (or not). Following the approach of (Lin, Snow, and Morgan 2011), for a topic t, we leverage a (small) set of user-curated topical hashtagsH to efficiently provide a large number of supervised topic labels for training. As standard for machine learning methods, we divide our training data into train and validation sets — the latter for hyperparameter tuning to control overfitting and ensure generalization to unseen data. As a critical insight for topical generalization where we view correct classification of tweets with previously unseen topical hashtags as a proxy for topical generalization, we do not simply split our data temporally into train, validation, and test sets and label both with all hashtags inH. Instead, we splitH into three disjoint sets H train, H t val, and H t test according to two time stamps t split and t test split for topic t and the first usage time Copyright c © 2017, Association for the Advancement of Artificial Intelligence (www.aaai.org). All rights reserved. #Unique Features From Hashtag Mention Location Term 95,547,198 11,183,410 411,341,569 58,601 20,234,728 Feature Usage in #Tweets Feature Max Avg Median Most frequent From 10,196 8.67 2 running status Hashtag 1,653,159 13.91 1 #retweet Mention 6,291 1.26 1 tweet all time Location 10,848,224 9,562.34 130 london Term 241,896,559 492.37 1 rt Feature Usage by #Users Hashtag 592,363 10.08 1 #retweet Mention 26,293 5.44 1 dimensionist Location 739,120 641.5 2 london Term 1,799,385 6,616.65 1 rt Feature Using #Hashtags From 18,167 2 0 daily astrodata Location 2,440,969 1,837.79 21 uk Table 1: Feature Statistics of our 829, 026, 458 tweet corpus. stamp htime∗ of each hashtag h ∈ H. In short, all hashtags h ∈ H with htime∗ < t split are used to generate positive labels in the training data, those with htime∗ ≥ t split are used for positive labels in the test data and the remainder are used for positive labels in the validation data. The key point to observe is that we not only partition the train, validation, and test data temporally, but we also divide the hashtag class labels temporally and label each data partition with an entirely disjoint set of topical hashtags. The purpose behind this training and validation data split and labeling is to ensure that learning hyperparameters are tuned so as to prevent overfitting and maximize generalization to unseen topical content (i.e., new hashtags). We remark that a classifier that simply memorizes training hashtags will fail to correctly classify the validation data except in cases where a tweet contains both a training and validation hashtag.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

A High-Performance Model based on Ensembles for Twitter Sentiment Classification

Background and Objectives: Twitter Sentiment Classification is one of the most popular fields in information retrieval and text mining. Millions of people of the world intensity use social networks like Twitter. It supports users to publish tweets to tell what they are thinking about topics. There are numerous web sites built on the Internet presenting Twitter. The user can enter a sentiment ta...

متن کامل

A Model for Detecting of Persian Rumors based on the Analysis of Contextual Features in the Content of Social Networks

The rumor is a collective attempt to interpret a vague but attractive situation by using the power of words. Therefore, identifying the rumor language can be helpful in identifying it. The previous research has focused more on the contextual information to reply tweets and less on the content features of the original rumor to address the rumor detection problem. Most of the studies have been in...

متن کامل

Improving Twitter Sentiment Classification Using Topic-Enriched Multi-Prototype Word Embeddings

It has been shown that learning distributed word representations is highly useful for Twitter sentiment classification. Most existing models rely on a single distributed representation for each word. This is problematic for sentiment classification because words are often polysemous and each word can contain different sentiment polarities under different topics. We address this issue by learnin...

متن کامل

Temporal Classification and Visualization of Topics in a Twitter Search Interface

Searching within Twitter is a challenging task; the short and cryptic nature of tweets leads to search results sets that may include information on many different topics. While many topic modelling approaches exist to extract the salient topics from the tweets, what is missing is a method for temporally classifying the topics and showing these to a searcher to help them understand the makeup of...

متن کامل

Examination of Emergency Medicine Physicians’ and Residents’ Twitter Activities During the First Days of the COVID-19 Outbreak

Introduction: Social media has become an important element of interaction and found itself a place in every aspect of our lives. This study examined the twitter activities of emergency medicine physicians and residents (EMP&R;) about the COVID-19 outbreak. Methods: The study concentrated on Twitter, a major social media network. To identify accounts owned ...

متن کامل

Examining the Automated Inference of Tweet Topics

The increasing volume of information exchange over online social networks (e.g. Twitter, Facebook) has led to the growing interest in technique for automated inference of the topic of individual posts/tweets in recent years. Short length, lack of a well defined set of topics, and use of acronyms in tweets are some of the reasons that make topic inference of tweets challenging. In this study, we...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:

دوره   شماره 

صفحات  -

تاریخ انتشار 2017